The Plant Phenome Journal
○ Wiley
Preprints posted in the last 30 days, ranked by how well they match The Plant Phenome Journal's content profile, based on 14 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Matuszynska, A.; Sansa, O.; Adekoya, F. J.; Akinyemi, O. O.; Anokye, E.; Bashir, O. B.; Boyny, Z. Z. F.; Chukwuka, M. K.; Corvest, E.; Dada, A. O.; DellAcqua, M.; Ehemba, G. L.; Finkbeiner, A. J.; Hamabwe, S.; Hodehou, D. A. T.; Kacheyo, O.; Kamfwa, K.; Mhango, K. J.; Abdullahi, W. M.; Munduwe, G.; Ntukidem, S.; Obisesan, O. K.; Odesina, I. S.; Ogechi, N.-U.; Olaoye, O. D.; Olayinka, M. M.; Osei-Bonsu, I.; Rilwan, K. O.; Stival, L.; Tehar, Z.; Tende, R. M.; To, J.; Ugochukwu, U. K.; Unger, A.; van Aalst, M.; Vrbic, D.; Zhang, C.; Theeuwen, T. P. J. M.; Kramer, D. M.; Kromdijk, J.
Show abstract
Photosynthesis is among the most consequential yet genetically complex traits in crop plants, and translating its natural variation into actionable genomic targets remains a central challenge for breeding climate-resilient varieties. To start addressing this, researchers are generating increasingly large, multi-environment field photosynthesis datasets. Yet, these data have been structurally under-analysed since their inception. Here we report the outcomes of the first dedicated hackathon focused on computational mining of such field data held in Accra, Ghana, in March 2026. Bringing together data scientists, plant physiologists, geneticists, and breeders from Europe and Africa, these interdisciplinary teams used photosynthetic data collected with hand-held fluorometers to genome-wide marker data across four crop species: cowpea (Vigna unguiculata), barley (Hordeum vulgare), common bean (Phaseolus vulgaris), and potato (Solanum tuberosum). Despite using different species and methods, independent teams identified the same three key findings. First, mechanism-informed feature engineering and dynamic modelling recover genetic signals that are not detected or discarded in standard analysis pipelines, resulting in traits with improved heritability and meaningful associations with yield. Secondly, machine learning methods proved effective at uncovering genetic associations, with temporally resolved features substantially outperforming single time-point measurements. Third, raw chlorophyll fluorescence and absorbance traces consistently contained more information and predictive power than the extracted parameters currently used. A defining feature of this event was having experimentalists and data scientists working together, enabling AI approaches to be grounded in domain knowledge and biological mechanisms rather than relying on data alone.
Harris, Z. N.; Braley, J.; Cassetta, E.; Crain, J.; DeHaan, L.; Van Tassel, D.; Miller, A.; Rubin, M. J.
Show abstract
Perennial grains represent a promising frontier for sustainable agriculture, but breeding progress is constrained by the accessibility of genotyping and the difficulty of evaluating complex traits expressed for multiple years after establishment across heterogeneous environments. Phenomic selection may help address these challenges by using inexpensive, scalable, high-dimensional phenotypes collected early in development, although the robustness of such predictions across breeding cycles remains uncertain. Here, we compared genomic selection and phenomic selection across two breeding cycles of Thinopyrum intermedium (intermediate wheatgrass; IWG; Kernza(R)), comprising approximately 2,280 individuals from maternal half-sib families evaluated across multiple field sites and years. We constructed relationship matrices from genomic markers and early-life stage phenomic data, including seed and leaf color (HSV), CropReporter multispectral reflectance and indices, and cycle-specific hyperspectral reflectance sensors. Genomic models provided the strongest predictions on average across all field traits in both cycles. Among phenomic predictors, leaf HSV was consistently the most informative, whereas CropReporter and hyperspectral data showed lower and more trait-dependent performance and seed HSV provided little predictive value. Genomic, leaf HSV, and CropReporter models transferred across breeding cycles with little apparent loss of predictive ability relative to within-cycle validation, demonstrating that their predictive signals were not restricted to a single breeding cycle. Early-life stage leaf HSV emerged as a practical, accessible tool for germplasm thinning and early-stage prioritization in perennial breeding programs. Despite limited similarity among relationship matrices, multi-relationship-matrix models rarely improved prediction beyond the stronger constituent single-relationship-matrix model. Together, these results show that early-life stage phenomic data provide reproducible information about agronomic performance expressed years later, but that predictor complexity and data integration do not guarantee improved prediction.
Stock, F.; Panda, S.; Poire, R.; Brown, T.; Akram, A.; Zheng, L.; Lei, H.; Zha, R.; Zhao, M.; Isabelle, S.; Martel, M.; Comeau, M.-A.; Hamel, L.-P.; Lavoie, P.-O.; D'Aoust, M. A.; Reithinger, H.; Saxena, P.; Stone, E. A.; Li, H.; Way, D. A.; Atkin, O. K.
Show abstract
Non-invasive, high-throughput phenotyping tools are needed that can identify environmental effects on plant structure and function to diagnose factors responsible for reduced growth in commercial and non-commercial settings. In this study, we explored whether the integration of 3D-multispectral (3D) and 2D-hyperspectral imaging (HSI), aided by machine learning (ML), could be used to identify environmental stress treatments imposed during plant growth. Controlled environment-grown Nicotiana Benthamiana plants were subjected to a range of abiotic treatments - including different growth irradiances, heat treatment and drought stress - with the treatments resulting in differences in shoot height, biomass, leaf area and spectral reflectance. ML models were trained to identify these treatments using morphological and spectral traits measured at 27, 29, 31, and 34 days after sowing (DAS). A 3D-multispectral scanner was used to obtain information on plant height, biomass, and leaf area. A visible and near-infrared (VNIR) HSI camera provided detailed spectral information for deriving spectral indices including the Normalised Difference Vegetation Index (NDVI), Photochemical Reflectance Index (PRI) and Normalized Difference Red Edge (NDRE). Manual measurements provided baseline comparative data. The 3D-multispectral scanner reliably estimated above-ground traits, with high correlations between manual and scanner-derived measurements. The ML models accurately differentiated among environmental stress treatments, with the fused 3D+HSI model achieving the best overall predictive performance across all evaluated metrics compared with models based on either imaging modality alone. Results demonstrated the effectiveness of combining 3D-multispectral and 2D-HSI data with ML analyses for non-destructive, high-throughput phenotyping. The integration of these techniques enabled non-destructive, high-throughput identification of environmental stress treatments imposed during plant growth.
Varela, S.; Ruhter, J.; Sacks, E.; Zheng, X.; Allen, D.; Hale, A.; Landry, C.; Kuang, X.; Long, B.; Zhu, Y.; Proma, S.; Kaur, S.; Jarquin, D.; Morrison, J.; Leakey, A.
Show abstract
The integration of digital technologies for high-throughput field phenotyping is critical for accelerating crop improvement in agriculture. However, extracting traits from remote sensing data remains constrained by fragmented workflows, manual intervention, and limited interoperability among existing tools, resulting in delays that hinder timely biological insight and decision-making. To address these challenges, we present PhenoStream (Phenotyping Streaming), a scalable, end-to-end cyberinfrastructure designed to automate the full lifecycle of aerial imagery-based phenotyping, from data acquisition to plot- and genotype-level inference. The framework integrates automated data ingestion from distributed field sites, geospatial processing, and AI-enabled trait extraction within a unified, user-accessible graphical interface. Its modular and extensible architecture supports adaptable trait modeling and seamless integration of new data sources, enabling deployment across diverse crops, environments, and experimental designs. We demonstrate the system across a large multi-location field trial network of bioenergy crops, where it enables high-throughput characterization of spatiotemporal growth dynamics, genotype-by-environment (GxE) interactions, and predictive modeling of key agronomic traits. By significantly reducing processing latency and manual effort, the platform facilitates near-real-time analysis and reproducible workflows. This work establishes a generalizable and scalable pathway for operationalizing very-high-spatial resolution aerial phenotyping in agricultural research. By bridging data acquisition and analytics, the end-to-end cyberinfrastructure provides a foundation for integrating heterogeneous and unstructured data streams--including remote sensing, environmental, and management data--toward data-driven decision making in agriculture.
Mejias, J.; Adreit, H.; Blanc, A.; Lubin, N.; Jolivet, C.; Guyot, V.; Brayle, O.; Poncelet, N.; Fournier, E.; Wicker, E. P.; Carlier, J.; Tharreau, D.; Ravel, S.
Show abstract
BackgroundThe quantification of fungal spores constitutes a fundamental metric in phytopathology, serving as the primary variable for inoculum standardization and being used as a proxy for disease severity. Historically, spore quantification has relied on manual hemocytometry, which remains the most precise counting process to date, where chambers such as the Malassez slide are used to count a subsample of the inoculum. However, this method applied manually is highly labor-intensive, time-consuming, and can be prone to operator-dependent variability. To overcome these limitations, we introduce MIRA (Microscopy Image Recognition & Analysis), a novel open-source software integrating You Only Look Once (YOLO) deep learning algorithms. Featuring a user-friendly graphical interface, MIRA is adaptable to multiple camera systems and supports advanced object detection models, including YOLOv11 and YOLOv26. ResultsWe demonstrate that MIRA can be used to accurately detect and count spores from several phytopathogenic fungi, automatically measure spore surface area, and to differentiate spores across different genera. In an exhaustive comparative analysis using Pyricularia oryzae spores as an example, MIRA was benchmarked against manual gold-standard counting slides (Malassez and Kova) and indirect spectrophotometric methods (SPARK). The P. oryzae model loaded via MIRA achieved a strong correlation (R = 0.96) with manual gold standards while reducing processing time by over 90% for high-concentration samples (10 spores/mL). Beyond this benchmark, we also successfully tested specific YOLO models designed to recognize macro- and microconidia of Fusarium oxysporum f. sp. cubense, a model for Pseudocercospora fijiensis, and a single multiclass model capable of identifying six different rice pathogenic fungi. We provide comprehensive tutorials for operating the software and training custom detection models for free using Roboflow and Google Colab. MIRA is available both as open-source Python code and as standalone executables for Windows and Linux. ConclusionsMIRA provides a rapid, accurate, and highly reproducible alternative to manual spore counting, effectively removing a major bottleneck in phytopathology workflows. By combining advanced YOLO-based deep learning with an accessible interface and comprehensive training resources, MIRA makes accessible automated image analysis for researchers without programming expertise. Moreover, MIRA drastically improves the efficiency of high-throughput disease phenotyping and can be adapted for a wide range of microscopic quantification tasks across various biological disciplines.
Nakata, R.; Hiraga, S.; Ishimoto, M.
Show abstract
Background and aims Plant volatile organic compounds (VOCs) change dynamically with plant development and in response to environmental conditions. However, their potential as non-invasive indicators of phenological progression remains poorly explored. In this study, we developed a framework integrating automated VOC sampling, time-resolved VOC profiling, and machine-learning analysis for the non-invasive assessment of plant phenology. Using soybean (Glycine max (L.) Merr.), we investigated whether development-associated temporal variation in VOC emissions could delineate and predict developmental phases. Methods We collected VOCs daily under controlled environmental conditions from 16 to 43 days after sowing, spanning the transition from vegetative to reproductive stages, using an automated sampling system coupled with thermal desorption-gas chromatograph-mass spectrometer (TD-GC-MS). To characterise temporal changes in VOC profiles associated with phenological progression, we analysed the daily VOC data using a multi-step pipeline combining statistical filtering and similarity-based network analysis. We defined VOC-derived developmental phases from similarity patterns in the VOC profiles, then developed and evaluated machine-learning models to predict these phases. Key results Seven VOCs exhibited distinct phase-dependent dynamics, including green leaf volatiles and monoterpenes showing characteristic temporal changes during phenological progression. Network-based clustering of VOC profiles resolved five developmental phases closely aligned with conventional developmental stages. A machine-learning model predicted these phases from the VOC profiles with high predictive accuracy on independent test data, demonstrating that phenological progression could be quantitatively inferred from VOC emission patterns. Conclusions Our findings support VOC profiling as a reliable and non-invasive approach for assessing phenological progression in soybean. By extracting temporally structured VOC signals, this framework captures developmental information that may be difficult to obtain through visual observation alone, particularly after canopy closure. VOC profiling offers a practical tool for monitoring crop developmental dynamics and has broader potential for plant phenotyping and precision crop management.
Castillo, M. P.; Oyebode, O. G.; Lenahan, A.; Orloski, A.; Wolfe, M.
Show abstract
White lupin (Lupinus albus L.) is a cool-season grain legume with seed crude protein of 33-47%, competitive with soybean (Glycine max L.) meal. It also fixes nitrogen and mobilizes soil phosphorus. Because soybean is a summer crop, white lupin can occupy Southeastern winter fields as a complementary protein source. Breeding for seed protein is limited by the cost and throughput of reference phenotyping. To determine how each is best deployed, we compared the utility of near-infrared spectroscopy (NIRS)-based phenomic selection with genomic selection based on 246,847 SNPs from low-pass, whole genome sequencing in a panel of Auburn University breeding lines and USDA National Plant Germplasm System germplasm. A handheld NIR calibration against Dumas reference protein reached screening-grade accuracy (R2 = 0.81). Under common cross-validation, phenomic predictive ability was 0.93 and genomic was 0.12. The low genomic value was consistent with moderate heritability (H2 = 0.33) and strong genotype-by-year interaction. Beyond predictive ability, NIRS recovered superior accessions the strictest selection intensity, and 40 to 60 reference assays sufficed to calibrate the model. Handheld NIRS is a low-cost tool for protein calibration and early-generation screening, while genomic prediction remains suited to parental selection, together supporting a complementary strategy for legume breeding Plain Language SummarySoybean meal is the main protein source for livestock and fish farms in the United States. Because soybean is a summer crop, many Southeastern fields sit idle or grow low-value cover crops in winter. White lupin, a cool-season legume whose seeds are as protein-rich as soybean meal, makes a good complementary winter crop: it yields high-protein grain while serving as a cover crop that fixes nitrogen and frees up soil phosphorus for later crops. In our early-stage lupin breeding program, measuring seed protein by standard lab methods is slow and costly. We built a calibration that lets a handheld scanner estimate protein from light, and compared it with predicting protein from the plants DNA. The scanner gave accurate, low-cost protein screening from only about 40-60 lab tests, while DNA-based prediction remains suited to guiding parent selection. Used together, these tools offer breeders a practical path to develop high-protein white lupin. Core ideasO_LIHandheld NIRS provides screening-grade prediction of white lupin seed crude protein. C_LIO_LISpectra carried more usable protein signal than markers by measuring seed chemistry directly. C_LIO_LINIRS and genomic prediction serve different stages of a white lupin breeding program. C_LIO_LIAbout 40 to 60 reference assays sufficed to calibrate NIRS to near-full accuracy. C_LI
Ewen, A.; Mendez, R. G.; Al-Shanoon, K.; Omoluabi, D.; Samarasinghe, A.; Oviedo-Ludena, M. A.; Huatatoca, K. C.; Glor, K.; Nabetani, K.; Kutcher, R.; Wang, L.; Stavness, I.; Jin, L.
Show abstract
Reliable and objective phenotyping is essential for plant breeding programs to characterize genetic variation and accelerate crop improvement. Conventional disease assessment relies on expert visual scoring, which is labor-intensive, subjective, and prone to inter- and intra-rater variability. Although image-based phenotyping methods have been proposed, many require manual intervention, specialized imaging setups, or single time-point measurements, limiting their ability to capture disease progression over time. Here, we present a pipeline for longitudinal plant disease phenotyping that quantifies wheat stripe rust and leaf rust progression from time-series images. The pipeline performs semi-automated leaf and automated pustule segmentation from images acquired in situ, enabling objective disease severity estimation with minimal user intervention and without requiring solid backgrounds or manual leaf manipulation or detachment. By extracting temporal traits, including disease severity trajectories and standardized area under the disease progress curve, the method provides a comprehensive characterization of disease development throughout infection. Association between automated and expert assessments was moderate for stripe rust (R2 = 0.58) and strong for leaf rust (R2 = 0.85), while expert inter-rater reliability was moderate for both diseases (ICC = 0.675 and 0.800, respectively). The proposed approach establishes a scalable and reproducible framework for longitudinal disease phenotyping in controlled environments, with broad applications in disease resistance screening and crop breeding.
Stutz, S. S.; Edquilang, R.; Bernacchi, C. J.; Ort, D. R.
Show abstract
Water-use efficiency (WUE), the ratio of accumulated plant biomass to water lost through transpiration has conventionally been determined using a destructive single-point measurement. Recent advances in high-throughput phenotyping now enable repeated, non-destructive estimation of biomass and WUE. However, these digital measurements must be statistically validated against conventional destructive methods to validate their use as reliable proxies. Therefore, we compared digital biomass determined point clouds produced from multispectral camera scanners with destructive harvests across eight harvests using Samsun tobacco grown under both drought and high-water conditions. WUE efficiency, calculated using the digital biomass estimated from a point cloud and gravimetric water use determinations, were compared to destructive harvest determinations. The coefficient of variation (CV) showed there were no significant differences in digital and destructive measurements for either biomass or WUE. Indicating that digital measurements can be used in place of destructive measurements. Drought plants used significantly less water and were significantly smaller than high-water plants from Harvests 4 through 8. However, there were no significant differences in the ratio of evapotranspiration to leaf area or WUE, indicating that drought plants were simply smaller and used less water than the high-water plants. This work validates that estimating plant biomass from a digital point coupled with continuous gravimetric determination of water use provides a reliable nondestructive measure of WUE in high-throughput measurements across the full plant life cycle.
Riaz, A.; Pearson, S.; Hunt, C.; Sukumaran, S.; Tao, Y.; Cooper, M.; Hammer, G.; Mace, E.; Jordan, D.
Show abstract
Tillering plasticity is a key adaptive trait in sorghum influencing resource use efficiency via a plants ability to adjust branching to neighbour density. Neighbour detection through red:far-red (R:FR) light sensing regulates this plasticity. While molecular pathways regulating tiller outgrowth are partly known, the genetic architecture underlying density-responsive tillering has not been resolved in any grass species. A sorghum diversity panel (n = 895) was evaluated over two growing seasons (2023 and 2024) with plant spacing ranging from 5 to 60 cm. A linear mixed model incorporating neighbour distance and tiller counts estimated genotype-specific response. GWAS was conducted on isolated plants (no neighbours within 60 cm) and on estimated responsiveness to neighbours. GWAS identified 52 baseline tillering QTLs and 50 for spacing responsiveness, with 10 overlapping, suggesting shared genetic control. Comparison with 41 R:FR pathway candidate genes revealed enrichment in responsiveness QTLs (5/50, 10%) versus baseline (0/52, 0%) (Fishers exact test, P = 0.025). Our model identified 40 unique density-responsive tillering QTL regions. Reducing genotype response to neighbour absence could be a selection target to develop water-efficient sorghum varieties where controlled architecture may be more valuable than natural plasticity.
Piao, X.; Lochocki, E. B.; McGrath, J.; Matthews, M. L.
Show abstract
Accurately modeling carbon (C) allocation is essential for predicting crop yield and the performance of new cultivars in various environments. Most crop models allocate C empirically, using fixed partitioning tables or harvest indices that prescribe allocation without representing the underlying physiology, limiting their predictive power under novel conditions. A mechanistic alternative, in which C allocation emerges from local utilization and transport, could instead respond dynamically to environmental changes, source-sink perturbations, and organ-level trait modifications. To achieve this design, we integrated a utilization-transport-resistance (UTR) allocation model into the Soybean-BioCro crop growth modeling framework. We calibrated and validated the model using organ biomass data from two soybean cultivars grown at two CO2 levels over eight seasons, achieving accuracy comparable to partitioning-based models while predicting more reasonable carbon allocation fractions. Further, the UTR-BioCro model predicted leaf and stem total nonstructural carbohydrate concentrations with reasonable accuracy compared to experimental measurements across the 2022 growing season. A local sensitivity analysis of the model parameters indicated that the onset of reproductive growth influenced yield more strongly than utilization or transport parameters suggesting the timing of this transition as a potential target for crop improvement. Finally, the UTR-BioCro model reproduced yield responses to source-sink perturbations including shading and pod removal, and captured the qualitative response to defoliation without requiring scenario-specific tuning as most partitioning approaches require. By grounding C allocation in physiological mechanisms, this work provides a foundation for predicting crop responses across diverse environments and engineered traits, supporting crop improvement for a changing environment.
Cazon, L. I.; Gonzalez, N. R.; Del Ponte, E. M.; Costa de Carvalho, A. C.; Asinari, F.; Camiletti, B. X.; Paredes, J. A.
Show abstract
Peanut smut, caused by Thecaphora frezzii, is an important constraint to peanut production in Argentina, but quantitative estimates of yield losses across environments remain limited. We quantified the relationship between disease incidence and kernel yield using 922 observations from 26 field studies conducted in Cordoba, Argentina, between 2021 and 2025. Study-specific incidence-yield relationships were analyzed using linear regression, random-effects meta-analysis, and linear mixed-effects models. Peanut smut incidence was consistently associated with yield reduction across studies. The estimated damage coefficient ranged from 24.2 to 28.7 kg ha-{superscript 1} per 1% increase in disease incidence, corresponding to a relative yield reduction of 0.74-0.87% of attainable yield. In contrast, attainable yield varied markedly among studies, ranging from 1,370 to 5,409 kg ha-{superscript 1}. Although an exploratory segmented analysis suggested a breakpoint near 12% incidence, subsequent moderator analyses, study- specific regressions, and normalized response curves provided no evidence of a biologically meaningful change in the damage coefficient across incidence or yield classes. These results indicate that differences among environments were primarily associated with attainable yield rather than with changes in the magnitude of disease-associated yield loss. The resulting damage function provides a quantitative basis for yield-loss assessment and disease management in peanut.
Zhao, J.; Ma, Y.
Show abstract
Germination percentage is an endpoint measure and therefore does not describe when an individual seed begins visible growth or how rapidly its radicle and plumule expand. We developed a time-resolved phenotyping workflow to quantify rice seed germination continuously in shallow-water culture. A single industrial camera moved along a 1 m rail and imaged three culture boxes at 1 h intervals for up to 80 h. The archive comprised 1,062 full-frame images and 6,372 seed-level repeated observations under the six-seed field-of-view configuration. A physical grid maintained seed identity through time and enabled individual regions of interest to be extracted. Whole-seed foregrounds were obtained with a pretrained U2-Net, and a masked RGB intensity rule separated newly emerging tissue from the darker hull. For each tracked seed, projected emerging-tissue area and interval growth rate were calculated. Three representative normally germinating seeds first showed measurable tissue at 48 h, yet subsequently followed distinct trajectories: final projected areas ranged from 2,605 to 4,700 pixels and peak interval growth rates ranged from 106.88 to 287.92 pixels h-1. B-1 accumulated 63.71% of its final visible area during 72-80 h, whereas B-3 accumulated 73.51% during 60-72 h. Thus, seeds with the same observed emergence interval can differ substantially in the timing and magnitude of post-emergence expansion. The workflow converts repeated images into biologically interpretable temporal phenotypes and provides a basis for nondestructive studies of rice seed vigor and germination heterogeneity.
Fukuda, H.; Sakamoto, T.; Yonemaru, J.-i.; Ogawa, D.
Show abstract
High temperature during grain filling increases rice grain chalkiness and deteriorates grain appearance under climate warming. Although several loci that reduce chalkiness have been identified, breeding strategies that integrate grain level heat tolerance with panicle level heat avoidance remain limited. Here we characterized SL2033, a chromosome segment substitution line carrying a long IR64 derived segment on chromosome 10, and evaluated the combination of the chromosome 10 segment with Appearance quality of brown rice 1 (Apq1), a quantitative trait locus associated with reduced heat induced chalkiness that acts at the grain level. Compared with its recurrent parent Koshihikari, SL2033 had longer flag leaves, altered vertical plant architecture, and lower panicle temperature. Total starch and protein contents were comparable between the two genotypes, whereas RNAseq analysis of the developing endosperm identified specific differences in heat, stress, and cell wall related transcripts. In a two year field trial, a pyramided line combining the SL2033 derived segment with Apq1 had the highest proportion of perfect grains and lowest frequencies of multiple chalky kernel types during the year with hotter grain filling conditions, with no detectable yield penalty. The pyramided line combined longer flag leaves, as in SL2033, with shorter panicle exsertion, as in an Apq1 near isogenic line, and had the lowest panicle temperature among the tested genotypes. Time series unmanned aerial vehicle imaging also detected genotype dependent differences in plant height during early grain filling, supporting distinct temporal patterns of plant development among the lines. These findings demonstrate that pyramiding genetic loci that confer panicle level and grain level heat tolerance is a promising strategy for improving rice grain appearance under high temperature field conditions, which are becoming increasingly prevalent.
Berlingeri, J. M.; Lo, S.; Riggs, M.; Yun, H.; Kamangir, H.; Ranario, E.; Uyehara, I. K.; Mayanja, I.; Lao, A.; Dramadri, I. O.; Ongom, P. O.; Boukar, O.; Palkovic, A.; Bailey, B. N.; Earles, J. M.; Huynh, B.-L.; Diepenbrock, C. H.
Show abstract
Cowpea (Vigna unguiculata [L.] Walp.) is a resilient grain legume and an important global source of dietary protein, yet the genetic and environmental basis of phenological and canopy development, as well as grain composition, remains incompletely characterized across production environments. In this study, we evaluated a cowpea multi-parent advanced generation intercross (MAGIC) population along an environmental gradient in California (with contrasting daylengths, temperatures, and soil types) using agronomic, grain compositional, and uncrewed aerial vehicle (UAV) and rover-enabled phenotyping. Near-infrared spectroscopy (NIRS) enabled assessment of grain compositional traits, while sensing-enabled time-series imaging captured canopy and reproductive dynamics. Quantitative trait locus (QTL) mapping identified 267 QTL, and genome-wide association studies (GWAS) detected 1,973 marker-trait associations. Integrating QTL mapping and GWAS results identified two major genomic hotspots affecting multiple traits. A chromosome 9 hotspot (5.8-6.0 Mb) was associated with flowering time and co-localized with sensing-enabled measures of flower and pod counts, plant height, and vegetation fraction, indicating broad effects on phenological and canopy development. A chromosome 8 hotspot (37.3-37.9 Mb) contained co-localized signals for seed weight, protein, starch, phytate, and moisture. A total of 22 prioritized candidate genes were identified within these and other loci with multi-environment QTL and GWAS support. Genomic predictive abilities were moderate to high for most traits and scenarios, with multi-trait MegaLMM outperforming RR-BLUP. Together, these results define major genomic regions controlling cowpea phenology, canopy development, and grain composition, and provide targets and strategies for breeding cowpea cultivars with favorable and environmentally resilient productivity and grain composition. Significance StatementTo dissect the genetic basis of cowpea productivity, adaptation, and grain composition, and how performance for these traits varies and can be predicted across environments, we combined multi-environment phenotyping, including sensing of canopy and reproductive traits, with quantitative genetic analyses in a multi-parental population. We identified genomic hotspots for seed size/composition and reproductive phenology and an across-environment predictive advantage for multi-trait vs. single-trait genomic prediction. Overall, these findings support the comprehensive improvement of cowpea.
Khan, F. S.; Yassin, A.; Rehman, S. u.; Sun, T.; Wang, X.; Sun, H.; Abe-Kanoh, N.; Su, Y. H.; Guo, L.; Ye, W.
Show abstract
Genome-wide association studies (GWAS) play a crucial role in unraveling the genetic foundations of complex traits in plants but are also hampered by the application of heterogeneous tools, incompatible file formats and disparate computational environments. Existing GWAS frameworks are often restricted to a single linear reference genome, limiting the capacity for the analysis of structural variations and presence/absence variations (PAV) within plant populations. These issues pose obstacles to reproducibility, scalability, and comprehensive investigations. Here, we present PlantOmicsGWAS, an open-source Python framework for reproducible plant genome-wide association analysis and genomic prediction. It integrates reference indexing, FASTQ quality control, alignment, variant calling, VCF normalization, PLINK conversion, linkage disequilibrium analysis, population-structure estimation, association testing, marker scoring, genomic prediction, and visualization within a unified Linux and HPC workflow. The framework supports conventional linear-reference analyses and includes an optional pangenome-oriented module for working with multiple assemblies and graph-derived variation. Using a Vitis benchmark dataset containing 120 accessions and 118,247 graph-derived variants, PlantOmicsGWAS reduced manual workflow fragmentation and generated standardized association outputs. This tool provides a modular and extensible platform for plant GWAS and pan-GWAS workflows while retaining compatibility with established command-line tools and common genotype formats. The GWAS workflow described herein is adaptable to a range of sequencing methods and plant genomes, bridging research on crop related issues across various biological levels, from the individual organism to entire populations. PlantOmicsGWAS implements Bayesian sparse linear mixed modeling (BSLMM) through GEMMA for multi-trait association discovery, while also supporting FaST-LMM, regression-based approaches, and machine-learning algorithms (Random Forest, XGBoost) as benchmarking alternatives. The PlantOmicsGWAS, a versatile toolkit is available at GitHub https://github.com/plantomicsgwas1-boop/PlantOmicsGwas_V1 and on Linux and HPC platform (https://pypi.org/project/PlantOmicsGwas/1.0.2/).
Ray, R.; Maloof, J.; Magney, T.
Show abstract
Leaf reflectance spectra are emerging as a viable substitute for gas-exchange measurements of photosynthetic capacity, with a community benchmark reporting that a spectrum accurately recovers most Farquhar-von Caemmerer-Berry (FvCB) parameters. This study re-scores the recovery under dataset-blocked, species-blocked, and leave-one-dataset-out designs, measuring the split-half reliability of each curated parameter. We constructed a convolutional encoder that maps a spectrum to the four parameters through a fixed, differentiable FvCB decoder trained on measured assimilation. A conspecific of 97.4% of held-out leaves were present in the training set, and accuracy is lost along the dataset axis but not along the species axis. Under blocked evaluation, a spectrum constrains a single capacity axis. Jmax25 retains only 17% of its recovery when Vcmax25 is held constant, and the Jmax25:Vcmax25 ratio is not predicted above a median null. The curated values of TPU25 are not reproducible, whereas those of Rday25 are well determined, but its recovery fails due to the loss. The published study measures interpolation rather than transfer, and spectra constrain less of the FvCB parameter space than assumed, including the carboxylation to electron transport balance. Routing predictions through explicit biochemistry makes identifiability measurable, although it does not improve prediction accuracy.
Pereira de Oliveira, L.; Attri, K.; Doran, L.; Leonelli, L. B.; Long, S. P.; Ainsworth, E.
Show abstract
Accelerating photoprotective regulation to improve carbon assimilation is a promising strategy to increase crop productivity. Although rapid non-photochemical quenching (NPQ) relaxation has been validated as a target through metabolic engineering, it remains unclear whether conventional breeding has improved this trait. Here, we investigated whether more than a century of soybean breeding enhanced NPQ relaxation alongside light-saturated carbon assimilation and seed traits. We evaluated a historical panel of 24 soybean genotypes across vegetative and reproductive developmental stages by integrating NPQ relaxation, gas exchange parameters, xanthophyll-cycle pigment profiles, expression of key photoprotective genes (VDE, PsbS, and ZEP), seed number and seed weight. NPQ relaxation parameters were not consistently associated with genotype release year, seed number, or seed weight at either developmental stage. The only exception was the amplitude of the rapidly relaxing NPQ component (AqE), which was negatively correlated with all three variables during the reproductive stage. In contrast, genotype release year was positively associated with maximum net CO2 assimilation rate (Amax), maximum carboxylation rate of Rubisco (Vcmax), maximum electron transport rate (Jmax), seed number, and seed weight, while Amax and Vcmax were positively correlated with seed number and seed weight. These findings indicate that the greater photosynthetic capacity of modern genotypes was not accompanied by faster photoprotective response. Thus, photoprotective regulation has not kept pace with gains in photosynthetic capacity under field conditions. We conclude that rapid NPQ relaxation remains an important target for synchronizing photoprotection with the high photosynthetic capacity of modern soybean lines.
Crawford, J. D.; Luebbert, C.; Baxter, I.; Schachtman, D.; Cousins, A. B.
Show abstract
A strategy to improve agricultural water productivity is to increase water use efficiency (WUE) at the level of plant transpiration through genetic selection. This requires detectable genetic variability in WUE and the ability to phenotype and select plants with higher WUE within a population. A proxy for phenotyping leaf level WUE by measuring carbon isotope signature ({delta}13Cleaf) has been supported by theory and data in C4 species. However, the functional relationship of {delta}13Cleaf and WUE in C4 species can be driven by genetics and environment. Therefore, a wide survey of existing natural variation is needed to quantify the heritability and identify various genetic factors that influence {delta}13Cleaf and WUE. In this study a genome-wide association panel was used to quantify the heritability of {delta}13Cleaf. We measured {delta}13Cleaf across a population of 360 genetically diverse lines of the C4 species Sorghum bicolor with single nucleotide polymorphic (SNP) markers determined from whole-genome resequencing. This analysis was conducted on two independent field environments where heritability of {delta}13Cleaf was evident and was driven by small genetic effects from loci that were consistently identified across environments. Candidate genes are presented that offer insights on future targets to manipulate and explore the functional relationship between {delta}13Cleaf and WUEi in C4 plants.
Dong, Y.; Li, J.; Li, F.; Luo, J.; Jia, Y.; Li, D.; Wang, L.; Su, X.; Hu, J.; Shang, Y.; Huang, S.; Zhu, Y.; Jia, Y.
Show abstract
Potato is an important non-cereal food crop worldwide. However, the limited number of functionally validated genes remains a major bottleneck to favorable allele stacking and genome design breeding in potato. Rapid advances in AI agents offer a promising means to support crop breeding by translating natural-language questions into coordinated data analysis and knowledge retrieval. Their reliable use for potato breeding, however, is constrained by fragmented multi-omics resources that lack consistent curation and machine-accessible interfaces. Here, we constructed an agent-ready potato multi-omics database integrating genomic resources from 150 potato accessions, 259 bulk RNA-seq samples, and 14 spatial transcriptomic datasets into a pangenome, a tissue expression atlas, co-expression networks, and spatial expression maps accessible through open APIs. We developed 39 potato-specific Agent Skills for reproducible bioinformatics analysis and comprehensive data and knowledge exploration, enabling natural-language questions to be translated into standardized data-retrieval and analysis tasks. By integrating direct evidence from potato studies, functions of homologous genes in Arabidopsis, rice, and maize, and tissue expression patterns, we generated genome-wide functional predictions for 37,658 genes in the DM reference genome. We further developed Potato Agent as a multi-user, browser-based platform with isolated workspaces and online result preview, reducing the technical burden of agent deployment and providing direct access to integrated data, knowledge, and workflows. Case studies demonstrated its capabilities in reproducible bioinformatics analysis, agent-assisted identification of a tuber development regulator, scientific data visualization, and haplotype-aware promoter analysis and sgRNA design. Together, the agent-ready database and Potato Agent provide an integrated infrastructure for functional gene discovery and hybrid breeding in potato.